Web Information Extraction Using Eupeptic Data in Web Tables

نویسندگان

Wolfgang Gatterbauer

Bernhard Krüpl

Wolfgang Holzinger

Marcus Herzog

چکیده

By leveraging on the redundant information on the Web, we are building a Web information extraction system that concentrates on eupeptic data in Web tables. We use the term eupeptic to describe such representations of information that allow for easy interpretation of the subject–predicate–object nature of individual data items. The system mimics a human approach to information gathering. It explicitly uses visual cues on rendered Web pages to locate tabular data; it uses keywords to identify relevant chunks of data that gets processed on a deeper level; and it expands its initial search to include more pages when it spots eupeptic data.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Web Information Extraction Using Eupeptic Data

متن کامل

Data Extraction using Content-Based Handles

In this paper, we present an approach and a visual tool, called HWrap (Handle Based Wrapper), for creating web wrappers to extract data records from web pages. In our approach, we mainly rely on the visible page content to identify data regions on a web page. In our extraction algorithm, we inspired by the way a human user scans the page content for specific data. In particular, we use text fea...

متن کامل

Presenting a method for extracting structured domain-dependent information from Farsi Web pages

Extracting structured information about entities from web texts is an important task in web mining, natural language processing, and information extraction. Information extraction is useful in many applications including search engines, question-answering systems, recommender systems, machine translation, etc. An information extraction system aims to identify the entities from the text and extr...

متن کامل

Automatic hidden-web table interpretation, conceptualization, and semantic annotation

The longstanding problem of automatic table interpretation still illudes us. Its solution would not only be an aid to table processing applications such as large volume table conversion, but would also be an aid in solving related problems such as information extraction, semantic annotation, and semi-structured data management. In this paper, we offer a solution for the common special case in w...

متن کامل

Schema Extraction for Tabular Data on the Web

Tabular data is an abundant source of information on the Web, but remains mostly isolated from the latter’s interconnections since tables lack links and computer-accessible descriptions of their structure. In other words, the schemas of these tables — attribute names, values, data types, etc. — are not explicitly stored as table metadata. Consequently, the structure that these tables contain is...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره شماره

صفحات -

تاریخ انتشار 2005

Web Information Extraction Using Eupeptic Data in Web Tables

نویسندگان

چکیده

منابع مشابه

Web Information Extraction Using Eupeptic Data

Data Extraction using Content-Based Handles

Presenting a method for extracting structured domain-dependent information from Farsi Web pages

Automatic hidden-web table interpretation, conceptualization, and semantic annotation

Schema Extraction for Tabular Data on the Web

عنوان ژورنال:

اشتراک گذاری